What am I looking at?
Each point is one gene: the identity vector the model learned for it during pretraining, reduced to two dimensions. Genes the model treats alike sit close together.
The model learned from gene expression across single-cell RNA-seq profiles, without sequences, family labels or pathways. Family annotations are overlaid afterwards, so you can see which known groups occupy similar parts of the learned space. Raise the detection threshold to focus on genes observed more consistently. Detection here is measured in validation-split cells.
The same 39 family categories are used across modalities. Unlabelled genes are shown in grey. Family membership does not guarantee a shared neighbourhood: search a gene or pathway and inspect its nearest neighbours. Shared expression structure is not proof of a physical interaction.
Hovera point to read gene, family and detection rate.
Clickto pin it and see its twenty nearest genes in the model’s space.
Familyswitch colouring to family and pick one to see it in pink against the rest.
Searcha gene (RPL3, KRT5) or a KEGG pathway (Ribosome, Oxidative phosphorylation) to light it up.